Papers with audio reasoning
Afrispeech Semantics: Evaluating Audio–Semantic Reasoning in Spoken Language Models Across Domains and Accents (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent multimodal models are trained on large collections of audio-text pairs using contrastive learning or nexttoken prediction objectives. |
| Approach: | They evaluate audio language models across five semantic and paralinguistic reasoning tasks: entailment, consistency, plausibility, accent drift, and accent restraint. |
| Outcome: | The evaluations assess models across five tasks including entailment, consistency, plausibility, accent drift, and accent restraint. |
Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in multimodal reasoning overlook the audio modality. |
| Approach: | They propose a large-scale audio language model for deep reasoning that leverages a multitask audio dataset. |
| Outcome: | The proposed model performs well across key benchmarks including MMAU-mini, AIR-Bench chat/foundation, and MELD. |
AUDITA: A New Dataset to Audit Humans vs. AI Skill at Audio QA (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing audio question answering benchmarks emphasize sound event classification or caption-grounded queries. |
| Approach: | They propose a large-scale, real-world audio question answering benchmark to evaluate audio reasoning beyond surface-level acoustic recognition. |
| Outcome: | The proposed model achieves 32.13% accuracy while demonstrating comprehension of audio . state-of-the-art models perform poorly, with average accuracy below 8.86%. |
Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent Large Audio Language Models (LALMs) have shown strong capabilities in audio understanding, yet their reasoning remains vulnerable to perceptual errors. |
| Approach: | They propose a large-scale dataset for **Perception-Aware Question Answering** that uses a hierarchical decoupling strategy to separate speech from environmental sounds and distinguishes among multiple speakers. |
| Outcome: | The proposed model improves on MMAU-mini, MMAR, and PAQA while maintaining comparable performance on multiple benchmarks. |